Papers by Gerard de Melo
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification using Graph Neural Networks? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Graph Neural Networks have been used for text classification, but only in domains with limited data characteristics. |
| Approach: | They compare graph representation methods for text classification using different architectures and setups. |
| Outcome: | The proposed graph representation methods outperform other models in document comprehension tasks. |
Model-Agnostic Bias Measurement in Link Prediction (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing work investigating social bias in factual knowledge graphs has focused on knowledge graph embeddings, so more recent classes of models achieving superior results by fine-tuning Transformers have not yet been investigated. |
| Approach: | They propose a model-agnostic approach for bias measurement leveraging fairness metrics to compare bias in knowledge graph embedding-based predictions (KG only) with models that use pre-trained, Transformer-based language models (KG+LM). |
| Outcome: | The proposed model-agnostic approach compares gender bias in occupation predictions with models that use pre-trained, Transformer-based language models (KG+LM). |
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual models have been released, but many of the world's languages are not covered. |
| Approach: | They propose a method that initializes the embedding matrix for a new tokenizer based on information in the source model's embeddable matrix. |
| Outcome: | The proposed method outperforms random initialization and previous work on language modeling and on a range of downstream tasks (NLI, QA, and NER). |
ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities (2023.findings-eacl)
Copied to clipboard
Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu
| Challenge: | a vision-language benchmark for human activity planning is designed for humans . the task is easy for humans, but challenging for SOTA deep learning models . |
| Approach: | They propose a vision-language benchmark for human activity planning that extends Charades with intents and builds on a multi-choice question test set. |
| Outcome: | The proposed benchmark evaluates the ability of systems to anticipate and plan human actions in a multimodal visionlanguage setting. |
Metaphor Suggestions based on a Semantic Metaphor Repository (L18-1)
Copied to clipboard
| Challenge: | Existing algorithms for suggesting metaphors have been used to find related words . corpus studies have found that metaphors are very pervasive even in formal language . |
| Approach: | They propose an algorithm that suggests metaphoric means of referring to concepts . they use MetaNet, a repository of conceptual metaphor, and lexical resources . |
| Outcome: | The proposed model expands the potential of the original repository by enabling new connections to be drawn. |
A Robust Self-Learning Framework for Cross-Lingual Text Classification (D19-1)
Copied to clipboard
| Challenge: | Recent advances in pretrained contextual representation models have made significant progress on a number of different English NLP tasks. |
| Approach: | They propose a robust framework to include unlabeled non-English samples in the fine-tuning process of pretrained multilingual representation models. |
| Outcome: | The proposed framework includes unlabeled non-English samples in the fine-tuning process of pretrained multilingual representation models. |
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation (P19-1)
Copied to clipboard
| Challenge: | Rhetoric is a vital element in modern Chinese poetry, and plays an essential role in improving its aesthetics. however, to date, it has not been considered in research on automatic poetry generation. |
| Approach: | They propose a rhetorically controlled encoder-decoder for modern Chinese poetry generation . their model captures various rhetorical patterns in an encoder and incorporates mixtures . |
| Outcome: | The proposed model outperforms state-of-the-art methods in terms of fluency, coherence, meaningfulness, and rhetorical aesthetics. |
Faithfully Explainable Recommendation via Neural Logic Reasoning (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing models for explainable recommendation have neglected faithfulness of KG reasoning . |
| Approach: | They propose to draw on interpretable logical rules to guide path-reasoning process for explanation generation. |
| Outcome: | The proposed method delivers high-quality recommendations and ascertains the faithfulness of the derived explanation. |
R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling (2021.acl-long)
Copied to clipboard
| Challenge: | Existing models with stacked layers do not explicitly model hierarchical structure of language understanding. |
| Approach: | They propose a recursive Transformer model based on differentiable CKY style binary trees to emulate hierarchical composition process. |
| Outcome: | The proposed model can predict words given their left and right abstraction nodes. |
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Numerical reasoning is ubiquitous in scientific research and financial analysis, but few benchmarks evaluate them by integrating numerical processing and mathematical reasoning. |
| Approach: | They propose a numerically-integrated hierarchical benchmark with 27,215 questions derived from 7,404 math word problems that spans 4 key cognitive aspects, 14 subcategories, and 2 modalities. |
| Outcome: | The proposed model improves Qwen-2.5 score with SOLVE and IRPO training. |
Correcting the Autocorrect: Context-Aware Typographical Error Correction via Training Data Augmentation (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that typographical errors are now ubiquitous . traditional spelling correction software is inadequate to correct typographical mistakes . |
| Approach: | They propose to generate typographical errors based on annotated spelling errors . they then use annotations to introduce errors into substantially larger corpora . |
| Outcome: | The proposed method generates typographical errors that require context-aware error detection . it also shows that machine learning can correct typographical mistakes based on the data . |
Assessing Emoji Use in Modern Text Processing Tools (2021.acl-long)
Copied to clipboard
| Challenge: | Emojis are textual elements that are encoded as characters but rendered as small digital images or icons that can be used to express an idea or emotion. |
| Approach: | They propose to use a set of popular NLP tools to assess the support of emojis in tweets. |
| Outcome: | The proposed methods show that many systems still have notable shortcomings when operating on text containing emojis. |
Improving Personalized Explanation Generation through Visualization (2022.acl-long)
Copied to clipboard
| Challenge: | Existing explainable recommendation models generate repetitive sentences for different items or empty sentences with insufficient details. |
| Approach: | They propose a visual-enhanced approach to generate rating scores and text explanations using visualization generation and text–image matching discrimination. |
| Outcome: | The proposed approach improves both the text quality and the diversity and explainability of the generated explanations. |
CITE: A Corpus of Image-Text Discourse Relations (N19-1)
Copied to clipboard
| Challenge: | a crowd-sourced resource characterizes inferences in image-text contexts in the domain of cooking recipes . a recent study has found that image-image presentations are more effective at integrating text and image . |
| Approach: | They propose a crowd-sourced resource for multimodal discourse characterizing inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations. |
| Outcome: | The proposed corpus enables a better understanding of communication and common-sense reasoning . it is particularly important for automating the understanding and generation of text-image presentations . |
Data Augmentation for Multiclass Utterance Classification – A Systematic Study (2020.coling-main)
Copied to clipboard
| Challenge: | a lack of sufficient training data for some categories can cause imbalanced data distributions . a weak classifier may miscategorize a request, resulting in customer dissatisfaction . |
| Approach: | They propose to use random resampling, word-level transformations and neural text generation to augment existing data to cope with imbalanced data. |
| Outcome: | The proposed methods improve utterance classification results by drawing on utterant variation. |
Context-Aware Interaction Network for Question Matching (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing models focus on word-level local matching and neglect the importance of contextual information. |
| Approach: | They propose a context-aware interaction network to properly align two sequences and infer their semantic relationship by using gate fusion layers. |
| Outcome: | The proposed model can accurately align two sequences and infer their semantic relationship on two question matching datasets. |
Assessing Combinational Generalization of Language Models in Biased Scenarios (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing work focuses on assessing in-domain knowledge, but shedding light on what pre-trained Language Models learn is important. |
| Approach: | They propose a method to assess a PLM's generalization capacity in biased scenarios by combining component combinations where it could be easy for the PLMs to learn shortcuts from the training corpus. |
| Outcome: | The proposed model can overcome distribution shifts in the training corpus and with sufficient data. |
Domain-Specific Sentiment Lexicons Induced from Labeled Documents (2020.coling-main)
Copied to clipboard
| Challenge: | Existing sentiment lexicons reflect abstract notion of polarity and do not do justice to substantial differences of word polarities between domains. |
| Approach: | They propose to use domain-specific sentiment lexicons to induce initial word intensity scores and train new deep models based on word vector representations to overcome the scarcity of the seed data. |
| Outcome: | The proposed models show that they perform well on review classification and cross-lingual word sentiment prediction. |
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations (2026.eacl-long)
Copied to clipboard
Yindong Wang, Martin Preiß, Margarita Bugueño, Jan Vincent Hoffbauer, Abdullatif Ghajar, Tolga Buz, Gerard de Melo
| Challenge: | Evaluating 9 state-of-the-art LLMs reveals two critical limitations: 61% of incorrect span predictions are semantically unrelated to actual errors. |
| Approach: | They propose a benchmark of 1,001 expert-annotated question-answer pairs with span-level error annotations derived from Reddit's r/AskScience. |
| Outcome: | Evaluating 9 state-of-the-art LLMs, we find that comparative judgment is paradoxically harder than independent detection when comparing answers side-by-side. |
InFact: Informativeness Alignment for Improved LLM Factuality (2025.findings-emnlp)
Copied to clipboard
| Challenge: | despite factual errors, LLMs tend to generate factual text that is factually correct but less informative than other, more informative choices. |
| Approach: | They propose an objective that prioritizes answers that are both correct and informative . |
| Outcome: | a new mechanism prioritizes correct and informative answers based on factual benchmarks . the proposed model improves both accuracy and factuality by maximizing the objective . |
Exploring Semantic Properties of Sentence Embeddings (P18-2)
Copied to clipboard
| Challenge: | Neural vector representations are ubiquitous throughout all subfields of natural language processing. |
| Approach: | They propose a framework that generates triplets of sentences to explore how changes in the syntactic structure or semantics of a given sentence affect their similarity. |
| Outcome: | The proposed framework generates triplets of sentences to explore how changes in the syntactic structure or semantics of a given sentence affect the similarities obtained between their embeddings. |
SLlama: Parameter-Efficient Language Model Architecture for Enhanced Linguistic Competence Under Strict Data Constraints (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large-scale language models (LLMs) have shown remarkable performance across a wide array of tasks. |
| Approach: | They propose an architecture that preserves parameter efficiency of tied models without sacrificing representational benefits of untied embeddings. |
| Outcome: | The proposed architecture achieves a 31.72% improvement in linguistic knowledge acquisition over the baseline model. |
Cross-Lingual Emotion Lexicon Induction using Representation Alignment in Low-Resource Settings (2020.coling-main)
Copied to clipboard
| Challenge: | Emotion lexicons provide information about associations between words and emotions. |
| Approach: | They use crowdsourcing to annotate words with Plutchik's 8 basic emotions, providing binary labels. |
| Outcome: | The proposed lexicons provide information about associations between words and emotions . the lexiconics are useful in emotional analyses of reviews, literary texts, and posts on social media . |
Query Distillation: BERT-based Distillation for Ensemble Ranking (2020.coling-industry)
Copied to clipboard
| Challenge: | Recent years have witnessed substantial progress in the development of neural ranking networks, but an increasingly heavy computational burden due to growing numbers of parameters and the adoption of model ensembles. |
| Approach: | They propose a two-stage distillation method that allows a smaller student model to be trained while benefiting from the better performance of the teacher model. |
| Outcome: | The proposed method shows higher-quality rankings compared to the teacher model. |
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit’s Showerthoughts (2024.starsem-1)
Copied to clipboard
| Challenge: | Recent Large Language Models (LLMs) have shown the ability to generate content that is difficult or impossible to distinguish from human writing. |
| Approach: | They compare GPT-2 and GPT-Neo fine-tuned on Reddit data and GTP-3.5 invoked in a zero-shot manner, against human-authored texts. |
| Outcome: | The proposed model can generate short, creative texts that are difficult to distinguish from human writing, but human evaluators rate them worse than the model. |
Generating Fine-Grained Open Vocabulary Entity Type Descriptions (P18-1)
Copied to clipboard
| Challenge: | Fig. 1 shows an example of a concise entity description presented to a user. |
| Approach: | They propose a dynamic memory-based network that generates a short open vocabulary description of an entity by leveraging induced fact embeddings and dynamic context. |
| Outcome: | The proposed network generates a short open vocabulary description of an entity . it can discern relevant information for more accurate generation of type description . |
Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to summarize dialogues are difficult due to insufficient training data and low information density. |
| Approach: | They propose a curriculum-based prompt learning method with self-training that gradually increases the degree of prompt perturbation, improving dialogue understanding and modeling capabilities. |
| Outcome: | The proposed model outperforms baseline models on the AMI and ICSI datasets and human evaluations show it is superior in the quality of the summary generation. |
FontLex: A Typographical Lexicon based on Affective Associations (L18-1)
Copied to clipboard
| Challenge: | a typographical lexicon provides associations between words and fonts . tens of thousands of fonts are available, and font choice affects perception of text, author and brand. |
| Approach: | They create a typographical lexicon providing associations between words and fonts by using affective evocations and word-emotion relationships. |
| Outcome: | The proposed typographical lexicon provides associations between words and fonts using affective evocations and word-emotion relationships. |
Data Augmentation with Adversarial Training for Cross-Lingual NLI (2021.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to train cross-lingual models with labeled data are subpar, resulting in subpar results. |
| Approach: | They propose a data augmentation strategy that enriches data to reflect more diversity in a semantically faithful way and leverages adversarial training regimens to achieve greater robustness. |
| Outcome: | The proposed approach improves cross-lingual inference by leveraging the data to reflect more diversity in a semantically faithful way. |
A Helping Hand: Transfer Learning for Deep Sentiment Analysis (P18-1)
Copied to clipboard
| Challenge: | Existing deep neural models for sentiment polarity classification require large amounts of training data. |
| Approach: | They propose to feed generic cues into the training process of deep convolutional neural networks for sentiment analysis. |
| Outcome: | The proposed approach improves sentiment polarity classification on a range of datasets in seven languages. |
Multi-Scale Distribution Deep Variational Autoencoder for Explanation Generation (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for generating explanations for recommender systems produce generic explanations that fail to incorporate user and item specific details. |
| Approach: | They propose a multi-scale distribution deepvariational autoencoder with a prior network that eliminates noise while retaining meaningful signals in the input. |
| Outcome: | The proposed models can generate explanations with concrete input-specific contents. |
EmoTag1200: Understanding the Association between Emojis and Emotions (2020.emnlp-main)
Copied to clipboard
| Challenge: | Emojis are increasingly used to convey affect, but their use is not trivial. |
| Approach: | They propose to use human-solicited association ratings to explore the connection between emojis and emotions to conduct experiments. |
| Outcome: | The proposed method can be inferred from word-level information when high-quality information is available. |
Inducing Universal Semantic Tag Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing semantic tags are useful for syntactically oriented downstream NLP tasks . but their size is limited and many words are out-of-vocabulary words . |
| Approach: | They propose to tagging words with semantic distinctions that are likely to be useful across semantic tasks. |
| Outcome: | The proposed semantic tagging scheme can predict unseen words with high accuracy . it distinguishes privative attributes from subsective ones, making it easier to discern fake detectives . |
ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing metrics for multimodal large language models only focus on token overlap and may not align with human judgment. |
| Approach: | They propose an open-source model that assesses the question answering abilities of multimodal large language models. |
| Outcome: | Experiments show that the ACE-M3 model performs better than existing models and is more reliable than existing metrics. |
PubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain? (2023.findings-eacl)
Copied to clipboard
| Challenge: | Medical visual question answering is a multimodal task that requires a system to understand both medical images and textual questions and infer associations between them. |
| Approach: | They propose a fine-tuned version of CLIP for the medical domain based on PubMed articles. |
| Outcome: | The proposed model improves accuracy up to 3% on two MedVQA benchmark datasets. |
Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text Representation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Chart-based models have shown great potential in unsupervised grammar induction, running recursively and hierarchically, but requiring O(n3) time-complexity. |
| Approach: | They propose a model-guided pruning method that scales to large language model pretraining by introducing a heuristic pruning method. |
| Outcome: | The proposed method significantly improves grammar induction quality and achieves competitive results in downstream tasks. |
Sentence Analogies: Linguistic Regularities in Sentence Embeddings (2020.coling-main)
Copied to clipboard
| Challenge: | Word vectors are often evaluated by assessing to what degree they exhibit regularities with regard to relationships considered in word analogies. |
| Approach: | They propose a number of schemes to induce evaluation data based on lexical analogy data as well as semantic relationships between sentences. |
| Outcome: | The proposed models reflect regularities in lexical analogies and semantic relationships between sentences. |
Interactive Question Clarification in Dialogue via Reinforcement Learning (2020.coling-industry)
Copied to clipboard
| Challenge: | ambiguous questions are a perennial problem in real-world dialogue systems. |
| Approach: | They propose a reinforcement model to clarify ambiguous questions by suggesting refinements of the original query. |
| Outcome: | The proposed model improves on real-world user clicks and shows significant improvements . it suggests that the original query is refined to clarify ambiguous questions . |
Guilt by Association: Emotion Intensities in Lexical Representations (2021.emnlp-main)
Copied to clipboard
| Challenge: | linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks . |
| Approach: | They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models. |
| Outcome: | The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data. |